Papers with situational harmful strategies

1 papers
Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation (2026.acl-long)

Copied to clipboard

Challenge: Existing static benchmarks for harmful content detection face limitations in scalability and diversity.
Approach: They propose a framework for synthesizing harmful content using persona-guided large language model agents.
Outcome: The proposed framework achieves a high success rate in harmful generation tests across multiple detection systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations